Skip to content

feat(health): add dependency-aware readiness and recovery diagnostics - #140

Open
woahwhattheheck wants to merge 9 commits into
RemitFlow:mainfrom
woahwhattheheck:latch/remitflow-134-readiness
Open

woahwhattheheck wants to merge 9 commits into
RemitFlow:mainfrom
woahwhattheheck:latch/remitflow-134-readiness

Conversation

@woahwhattheheck

@woahwhattheheck woahwhattheheck commented Sep 24, 2026 •

Copy link
Copy Markdown

Summary

Closes #134.

Separates process liveness from traffic readiness and adds bounded, redacted dependency probes so a process can no longer report healthy while its store, payment provider, or FX dependency is down.

Current behavior

  • GET /api/health/live is process-only and dependency-free, and is routed ahead of the business API rate limiter so dependency outages or exhausted business quotas do not make the liveness probe flap.
  • GET /api/health/ready probes store, payments (Stellar), and FX under a normalized per-check deadline (HEALTH_CHECK_TIMEOUT_MS, default 1000ms).
  • Failures return HTTP 503 with stable dependency-scoped reason codes such as STORE_UNAVAILABLE and PAYMENTS_TIMEOUT; raw provider messages, stacks, credentials, and connection strings are not returned.
  • Completed checks are re-evaluated on the next request, so ordinary outage recovery does not require a restart.
  • Overlapping readiness requests share one unfinished operation per dependency/adapter generation. Each caller keeps its own deadline, preventing repeated polls from multiplying hung provider work.
  • Adapter replacement starts a new generation immediately; a late predecessor completion cannot evict the replacement.
  • Timer inputs are normalized to Node's supported delay range so overflow or infinite values cannot become accidental 1ms deadlines.
  • Payment/FX adapters may resolve synchronously or asynchronously; immediate and post-deadline rejections stay on the redacted readiness failure path.

Design boundaries

  • Deadlines bound waiting, not arbitrary provider execution; a plain Promise cannot be cancelled generically.
  • A truly never-settling unchanged adapter leaves one shared unfinished operation rather than spawning one operation per readiness poll. Recovery occurs when it settles or the adapter is replaced.
  • The shipped payment/FX adapters are mocks/configuration checks, not live provider reachability probes. Real transports should also enforce their own cancellation/timeouts.
  • /api/health/ready remains behind the normal API middleware/rate-limit path; only liveness bypasses the business quota.
  • No broad transfer/user-flow rewrite is included.

Acceptance mapping

Criterion Current implementation
Liveness remains responsive during outages process-only /api/health/live path, ahead of business quota
Readiness changes and recovers fresh dependency evaluation after completed checks; adapter-generation replacement recovery
Dependency checks cannot hang the readiness response normalized per-caller deadlines
Repeated timeouts do not multiply provider work one shared unfinished operation per dependency/adapter generation
Redacted diagnostics fixed dependency-scoped reason-code taxonomy
Status behavior ready => 200; dependency failure/timeout => 503
Original failure regression store/payment/FX unavailable or timed out => not_ready while liveness remains available

Validation boundary

Exact current head: 364882b6ace3336480f7a358eb20283bdc61a40d (9 commits / 17 files).

Historical focused evidence retained in-repo:

  • probe-sharing repair at b4c9505ee1ac758ebb3fff659bbe39747febaa6b: 29 focused tests passed in run 37193635966, including repeated timeout sharing and recovery
  • timer-range hardening: the maintained three-case bounded timer regression passed against its changed source and is documented in docs/READINESS_TIMER_RANGE.md

The current exact head also adds dependency-scoped reason-code coverage. Its hosted CI run 37211797983 is action_required; there are no current-head status contexts or review submissions. Therefore this PR does not claim a fresh current-head full-suite or CI-green result.

Compatibility

/api/health and /api/health/live keep their existing process-health semantics. /api/health/ready exposes structured per-dependency entries (status, latencyMs, optional reason) plus the effective timeout budget.

Separate process liveness from traffic readiness. Ready probes now check
store, payments (Stellar), and FX under a per-dependency timeout, return
redacted reason codes on failure, and re-evaluate every request so recovery
does not require a restart. Liveness stays dependency-free so outages do not
flap orchestrator restarts.

Closes RemitFlow#134
Validate the actual loaded rate values for every advertised currency so a
missing or corrupt FX table cannot report ready while rate requests fail.
Keep the existing success payload, FX_UNAVAILABLE code and corridors.

Exercise table loss, liveness and same-process recovery through the actual
HTTP app, plus partial invalid-rate states. The new data-loss regression
failed against the previous source with 200 instead of 503.

Validation: npm test passed all 271 tests; no tests skipped.
Serve only GET/HEAD liveness before the global API budget, preserving common middleware and the existing health controller. Readiness, business traffic and unmatched routes stay limited.

Native createApp before/after: 17 local HTTP requests per phase prove availability after quota exhaustion and that liveness polls do not spend the business budget. Existing FX outage/recovery behavior is preserved.

All 273 tests from the unchanged npm test selection pass on Node 24.19.0 with --test-concurrency=1 and retained packages matching all 76 lockfile versions. Focused health checks: 17 pass; exact parent with identical final tests: 2 fail / 15 pass. Syntax and whitespace checks pass. CI uses Node 22 and remains a separate hosted gate; services retain their in-memory/mock boundary.
@woahwhattheheck

Copy link
Copy Markdown
Author

Current-head QA at 364882b6ace3336480f7a358eb20283bdc61a40d: #134’s liveness/readiness split, bounded per-caller deadlines, redacted dependency reason codes, recovery after completed failures, and shared in-flight probe generation handling are present. Liveness bypasses business rate limiting while readiness remains dependency-aware. The uncancellable-promise boundary is documented: an unchanged never-settling adapter is shared rather than multiplied until it settles or is replaced. I did not run a local suite. GitHub has no current-head status contexts; CI run 37211797983 is action_required, not a test failure. Source unchanged.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

feat(backend): add dependency-aware readiness and recovery diagnostics

1 participant